Back

ACS Catalysis

American Chemical Society (ACS)

Preprints posted in the last 30 days, ranked by how well they match ACS Catalysis's content profile, based on 18 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit.

1
CatESO: Differentiable Enzyme Sequence Optimization Guided by Substrate-Aware kcat Prediction

Gan, Z.; Xu, Y.; Xu, J.; Wu, Z.; Huang, J.; Yin, J.; Chen, G.; Zhang, J. Z. H.

2026-07-06 biochemistry 10.64898/2026.07.04.736506 medRxiv
Top 0.1%
9.6%
Show abstract

Enzymes drive biological chemistry and offer greener routes to chemicals, materials and medicines, yet their broader use as biocatalysts is often limited by insufficient catalytic turnover. Improving turnover is hard: measured rate constants are scarce and protein sequence space is vast. Deep-learning models now predict the turnover number, Kcat, with growing accuracy, but they are typically applied after sequence generation to score or filter candidates, which separates the kinetic objective from the design itself. To bridge the gap between sequence generation and kinetic evaluation, we introduce CatESO, a differentiable sequence optimizer that enables direct, gradient-guided design of substrate-specific catalytic turnover. By backpropagating through a cross-modal Kcat predictor under continuous sequence relaxation, CatESO co-optimizes predicted catalytic activity, evolutionary plausibility and structural integrity in one end-to-end framework, using ESM-2 and ESMFold to keep designs evolutionarily plausible and foldable. Across seven stringent out-of-distribution enzymes spanning EC classes 1-7, CatESO raised model-predicted Kcat for the vast majority of designs, with a median predicted fold change of 1.52 while every variant retained a pLDDT above 70. Against RFdiffusion3-LigandMPNN pipeline and ZymCtrl, CatESO struck a better balance between predicted activity and structural confidence. By making substrate-conditioned kinetic objectives differentiable, CatESO carries differentiable protein design beyond structure- and binding-centred goals to enzyme catalytic function, giving a general route to function-oriented enzyme engineering.

2
Targeted mining of plastic-associated metagenomes uncovers a novel thermostable PETase expanding scaffold space for engineering

Rigkos, K.; Bezantakou, D.; Antoniadis, K.; Antonopoulou, I.; Zarafeta, D.; Skretas, G.

2026-07-10 biochemistry 10.64898/2026.07.10.737215 medRxiv
Top 0.1%
7.9%
Show abstract

Enzymatic depolymerization of polyethylene terephthalate (PET) has advanced rapidly, alongside a growing volume of publicly available metagenomic data from microbial communities under sustained selective pressure from plastic exposure. Reasoning that such environments may harbor underexplored polyester-active enzymes, we developed a targeted mining workflow that screens exclusively plastic-associated datasets through multi-step bioinformatic filtering--integrating catalytic-motif screening, disulfide-topology validation, structural-similarity scoring, and phylogenetic profiling--to recover high-confidence PETase candidates. Applied to 271 plastic-associated metagenomes, the pipeline yielded 21 non-redundant candidates, several of which combine the Type I catalytic motif (GHSMGGGG) with Type II-like extended loops and secondary disulfide bonds. Two candidates were experimentally confirmed as PET hydrolases; the more active, PET-KR1, is a thermostable enzyme (Tm = 66.5 {degrees}C) that depolymerizes PET across a broad temperature range, with markedly higher productivity on powdered than on film substrate. PET-KR1 achieved optimal depolymerization at 50 {degrees}C, yet at 60-65 {degrees}C, where total yields declined, the product pool was more strongly enriched in the terminal monomer TPA, suggesting that thermostability and substrate accessibility are the primary targets for further engineering. Molecular dynamics simulations revealed a conserved hydrophobic binding network around the catalytic serine, consistent with established PETase substrate-recognition modes, and rational disulfide engineering raised the melting temperature by 3.5 {degrees}C, confirming amenability to further optimization. Overall, PET-KR1 expands the scaffold space available for PETase engineering, while the discovery workflow, built entirely on publicly available tools and open-access data, provides a reproducible strategy for metagenomic mining of novel PET-degrading enzymes toward biocatalytic PET recycling.

3
Directed evolution of the Fe-nitrogenase for CO2 reduction to hydrocarbons

Oehlmann, N. N.; Schmidt, F. V.; Chen, J.; Prinz, S.; Zarzycki, J.; Claus, P.; Kahnt, J.; Erb, T. J.; Rebelein, J. G.

2026-07-09 biochemistry 10.64898/2026.07.08.737278 medRxiv
Top 0.1%
7.2%
Show abstract

The iron (Fe) nitrogenase drives bacterial methane (CH4) formation by converting carbon dioxide (CO2) to CH4 in a single enzymatic step. Enhancing the initial CH4 formation activity of Fe-nitrogenase and expanding the product spectrum to hydrocarbon chains could lead to a route for sustainable feedstock chemicals. Here, we performed the first directed evolution campaign on the Fe-nitrogenase aimed at optimizing the hydrocarbon production. We achieved an ~8-fold increase in CH4 formation by Fe-nitrogenase expressing Rhodobacter capsulatus cultures in three rounds of site-saturation mutagenesis. The best performing mutant (F362ManfD, Y85FanfD, T360SanfD) extends the in vivo product spectrum of the nitrogenase to ethane (C2H6) and exhibits 6-fold higher rates for CO production in vitro, whereas the formation of the undesirable byproduct formate was abolished. Electron microscopy-based structural analysis identified a methionine and water potentially stabilizing the transition state and fine-tuning the CO2 reduction mechanism and activity.

4
Substrate-dependent epistasis probes active site intramolecular wiring

Buda, K.; Miton, C. M.; Vogt, C.; Tokuriki, N.

2026-07-03 biochemistry 10.64898/2026.07.02.736193 medRxiv
Top 0.1%
6.4%
Show abstract

Enzyme adaptation toward novel substrates involves the rewiring of intramolecular residue networks, yet how this rewiring differs across multiple substrates, and how it underpins functional trade-offs and promiscuity, remains poorly understood. Here, we profile all 64 combinations of six key mutations in a phosphotriesterase across nine structurally diverse substrates spanning three chemical classes (organophosphates, esters, and lactones), thus generating a multi-dimensional map of epistasis and promiscuity within the phosphotriesterase's active site. We developed a statistically robust reference-based analysis pipeline incorporating error propagation and significance testing to move beyond global epistatic trends and resolve idiosyncratic, substrate-dependent intramolecular wiring in specific genetic backgrounds. Simulations confirm that this pipeline reliably identifies genuine higher-order epistatic interactions while minimizing false positives. We reveal that intramolecular network wiring varies substantially between substrates, even within the same chemical class, with notable divergences between the adaptive target substrate 2-naphthyl hexanoate and its shorter-chain ester analogs. Key higher-order networks, including d233E/h254R/l271F and l271F/f306I/i313F, exhibit substrate-specific epistatic signatures that discriminate between subtle structural features such as acyl chain length, leaving group identity, and heteroatom substitution. These substrate-dependent rewiring events account for observed functional trade-offs, particularly the strong anti-correlation between the adaptive and native substrates. Collectively, these findings demonstrate that comprehensive cross-substrate epistatic profiling, paired with rigorous statistical analysis, provides a powerful framework for dissecting the molecular basis of enzyme promiscuity and the trade-offs that define adaptive evolution.

5
MAERM: Predicting Enzyme-Reaction Matching Relationships with a Mixed-Attention Model

Liu, T.; Zhai, S.; Lin, S.; Zhan, X.; Deng, J.; Liu, H.; Siu, S. W. I.

2026-07-10 bioinformatics 10.64898/2026.07.06.736902 medRxiv
Top 0.1%
5.4%
Show abstract

Harnessing enzyme specificity requires a thorough understanding of enzyme promiscuity, which determines enzymes catalytic scope; however, measuring this scope still relies heavily on labor-intensive analytical approaches. While data-driven approaches have emerged to predict the catalytic scope of enzymes, these methods continue to face challenges such as restricted datasets and insufficient integration of enzyme structural information and reaction transformations. Here, we introduce MAERM, an innovative mixed-attention model designed to predict enzyme-reaction matching relationships. Built on our MAERM-DB, a dataset with broad coverage of validated and chemoenzymatic catalysis data, MAERM utilizes a local-global attention module to integrate multimodal enzyme information with fine-grained reaction representations, thereby predicting enzyme-reaction matching probabilities. Results show that MAERM consistently outperforms all baselines, with an average F1-score of 0.984. Notably, on challenging test samples with less than 40% sequence identity to the training set, MAERM outperforms the second-ranked model by 5.9% in F1-score. In addition, MAERM achieves the highest top-10 success rate of 51.7% on Enzyme-405 and the highest balanced accuracy of 0.697 on BioCat-547, further supporting its generalizability in enzyme screening and chemoenzymatic catalysis. Finally, MAERM can serve as an efficient scoring module. When integrated with ProteinMPNN, MAERM has successfully guided novel enzyme design for two carbonyl reduction reactions, resulting in enhanced catalytic potential for the native substrate and demonstrating broad compatibility. Overall, MAERM has the potential to reduce the experimental cost of measuring enzymes catalytic scope, facilitate enzyme design, and ultimately accelerate the design-build-test-learn cycle in enzyme engineering.

6
Unraveling Xanthomonas acetyltransferase GumG: decoding catalytic promiscuity, understanding the mechanism, and enhancing enzymatic versatility

Liu, Y.; Ruehmann, B.; Melse, O.; Bayaraa, T.; Kampl, L.; Doering, M.; Sieber, V.

2026-06-23 microbiology 10.64898/2026.06.22.733816 medRxiv
Top 0.1%
4.8%
Show abstract

Xanthan is a structurally complex exopolysaccharide produced by Xanthomonas campestris and one of the most extensively studied microbial biopolymers. As a sustainable alternative to petroleum-based polymers, its broader application requires precise control of polysaccharide decoration, yet the enzymatic basis of these modifications remains incompletely understood. Here, we characterise the activity and substrate scope of GumG, an AT-3 domain-containing membrane-bound acetyltransferase responsible for xanthan O-acetylation. Using mass spectrometry in combination with in vitro and in vivo assays, we show that GumG mediates non-specific acetylation of the outer mannose residue and displays pronounced substrate promiscuity. GumG also exhibits limited propionyltransferase activity, enabling the biosynthesis of hybrid acetylated-propionylated xanthan at an 8.27:1 ratio. Molecular docking and analysis of 31 xanthan variants identify a cytoplasmic substrate-binding pocket defined by Val67 and Phe71 that governs donor specificity, and an engineered GumG variant (F71L) shows enhanced propionyltransferase activity. In addition, a periplasmic His40-Trp143-Asp246-His297 motif is proposed to constitute the catalytic center. Together, these findings provide mechanistic insight into GumG multifunctionality and establish a framework for engineering xanthan derivatives with tailored physicochemical properties.

7
Integrating Machine-learning and Ultra-high-throughput Screening for Enzyme spaces exploration

Ke, Y.; Zhang, Y.; Fang, M.; Zhao, J.; Zhu, H.; Xu, Z.; Cao, L.

2026-06-24 biochemistry 10.64898/2026.06.23.733994 medRxiv
Top 0.1%
4.2%
Show abstract

AbstractThe systematic navigation of biocatalyst space is constrained by elusive structure-activity rules and a lack of evolutionary history. Here, we present IMUSE, a strategy integrating machine learning with ultra-high-throughput screening. By screening millions of droplet-encapsulated de novo enzymes, we generated massive synthetic sequence-structure datasets to train models that capture their complex fitness landscapes and biophysical principles. These models effectively guide functional exploration across both sequence and novel structure spaces. IMUSE identified synergistic triple mutations yielding [~]5-fold activity improvements and discovered active second-generation designs with novel catalytic pockets, boosting the experimental success rate >4.9-fold ([~]30%). This work demonstrates how synthetic fitness landscapes bridge the data gap in de novo enzyme space, transforming stochastic search into deterministic navigation to unlock highly proficient biocatalysts beyond natural boundaries.

8
Discovery and structural analysis of glycoside hydrolase family 176 α-1,2 glucosidase from Arthrobacter humicola A8F5

Yasukochi, R.; Suzuki, T.; Toraya, T.; Hino, K.; Mori, T.; Kashima, T.; Miyanaga, A.; Watanabe, H.; Fushinobu, S.

2026-07-03 biochemistry 10.64898/2026.07.01.735942 medRxiv
Top 0.1%
3.9%
Show abstract

Glycoside hydrolases (GHs) exhibit remarkable specificity dictated by the structural configuration of their target glycosidic linkages. While enzymes that process -1,4- and -1,6-linkages in starch or glycogen are well-characterized, those acting on less common bonds, such as -1,2-glucosidic linkages, remain largely underexplored. In this study, we report the discovery and structural elucidation of a novel -1,2-glucosidase from Arthrobacter humicola A8F5 (A8F5 glucosidase), representing a newly uncovered activity within the poorly characterized GH176 family. Biochemical characterizations revealed that A8F5 glucosidase exclusively cleaves -1,2-linkages via an anomer-inverting mechanism, with a distinct preference for short kojioligosaccharides. To circumvent crystallization obstacles caused by high loop flexibility and translational non-crystallographic symmetry, we engineered a loop-truncated variant. This strategy enabled the determination of high-resolution (up to 1.79 [A]) crystal structures of the enzyme in its ligand-free form and in complex with kojibiose, kojitriose, and selaginose. A8F5 glucosidase adopts a (/{beta})6-barrel fold characteristic of clan GH-G. Complementing the crystal structures with AlphaFold3 prediction demonstrated that two prominent active-site loops (loops 3 and 4) adopt a closed conformation that constricts the catalytic pocket, rendering the architecture suitable for short oligosaccharide recognition while restricting access to larger polymers. Furthermore, sequence similarity network analysis highlights vast, uncharacterized functional diversity within the GH176 family. These findings revealed that the GH176 enzyme recognizes and hydrolyses -1,2-glucosidic bonds through a structural framework distinct from that of the previously known clan GH-L GH65 kojibiose hydrolase, expanding the known functional landscape of this enzyme group toward rare -glucans.

9
Fragment Based Active Site Exploration of Urethane Hydrolases Reveals a Diversity of Urethane Binding Modes

Bicer, D.; Kochubei, D.; Graham, R.; Pena-Diaz, S.; Rotilio, L.; Villadsen, N. L.; Sommerfeldt, A.; Johansen, M. B.; Sandahl, A.; Thirup, S. S.; Morth, J. P.; Otzen, D. E.

2026-07-07 biochemistry 10.64898/2026.07.06.734427 medRxiv
Top 0.1%
3.5%
Show abstract

Recent advances in the discovery, characterisation, and engineering of urethanases provide new opportunities for the sustainable biocatalytic degradation of polyurethane waste. A mechanistic understanding of enzyme-plastic interactions is essential for structure-based engineering to enhance urethanase activity. However, the extremely complex and hydrophobic nature of polyurethane makes it challenging to elucidate the structural basis of enzyme-plastic interactions. Here, we used a fragment-based approach to characterise the active sites of two novel urethanases with different catalytic scaffolds, employing both a crystallographic fragment-screening (FASE) campaign and soluble fragments of plastic-like analogues that mimic the substrate, transition state, or product. FASE identified new substrate-binding subpockets while interactions of plastic mimetics in the active site provided a mechanistic understanding of the recognition and binding of polyurethane fragments by these subpockets. These results highlight a diversity of binding modes among urethanases toward different polyurethane fragments.

10
Function-guided design of active enzymes

Hu, M.; Wu, L.; Yang, Y.; Li, F.; Zhu, L.

2026-06-29 bioinformatics 10.64898/2026.06.27.735025 medRxiv
Top 0.1%
3.2%
Show abstract

Designing enzymes from functional descriptions remains challenging because catalytic activity is governed by sequence-structure-function relationships. Here we present EnzymeArt, a function-conditioned enzyme-design framework centred on a generative sequence model. EnzymeArt couples function-conditioned sequence generation with structure-guided refinement, annotation checks and substrate-aware computational prioritization to select candidates for synthesis and biochemical testing. Across alcohol dehydrogenase (ADH), malate dehydrogenase (MDH) and triacylglycerol lipase design campaigns, 57 of 60 synthesized designs showed crude-lysate activity above matched background controls. Purified representatives further showed quantitative steady-state catalytic activity. The best designed ADH reached kcat = 223.7/s and exceeded a wild-type reference under matched conditions, an MDH reached kcat = 267.57/s despite having only 33% sequence identity to its closest BLASTP hit, and a designed lipase hydrolysed both short- and long-chain triglycerides with apparent activity modestly above that of a commercial lipase reference. Together, these results establish a route for converting functional descriptions into experimentally validated enzyme designs with quantitative steady-state kinetic activity.

11
Influence of Primary Coordination Sphere on Anion Rebound Selectivity in Nonheme Fe Enzyme-Catalyzed C(sp3)-H Functionalization: A Comparative Experimental and Computational Study of EgtB and ACCO

Yang, Y.; Zhao, L.; Guo, R.; Mai, B. K.; Chen, H.; Liu, P.

2026-07-13 biochemistry 10.64898/2026.07.10.737789 medRxiv
Top 0.1%
3.1%
Show abstract

Developing enzymatic mechanisms for C-F bond formation remains a long-standing challenge. Here, we repurposed the biosynthetic nonheme Fe enzyme EgtB, which features a three-histidine facial triad, to catalyze C(sp3)-H fluorination reactions. Directed evolution of EgtB afforded two new-to-nature fluorine atom transferases with opposite enantiopreference, EgtBCHF1 and EgtBCHF2, with up to 28-fold improved total activity. In contrast to our previously evolved nonheme Fe fluorine atom transfer biocatalyst ACCOCHF, which contains a two-histidine-one-carboxylate facial triad, the evolved EgtBCHF variants displayed unexpected hydroxylation activity. 18O-labeling experiments showed that the hydroxy group originated from water rather than residual O2. Computational studies suggested that the three-histidine-supported Fe(III) center exhibits enhanced Lewis acidity compared to the two-histidine-one-carboxylate system, allowing deprotonation of Fe(III)-bound water to form a Fe(III)-OH species to catalyze radical hydroxylation. Primary coordination-sphere mutagenesis in EgtB and ACCO further supported the critical role of Fe coordination chemistry in controlling radical rebound reactivity and selectivity. Computational studies revealed that Fe coordination chemistry strongly influences both fluorine atom abstraction and radical rebound, with the intrinsic C-X (X = F, OH, and N3) bond forming radical rebound preference following the order N3 > OH > F. Furthermore, multivariate linear regression analysis revealed that fluorine atom abstraction is primarily governed by the intrinsic Fe-F bond strength, whereas fluorine rebound is predominantly controlled by the electronic structure of the Fe(III) intermediate. Together, these findings provide mechanistic insights into nonheme Fe enzymology and reprogramming toward selective radical rebound reactions, including challenging C-H fluorination. Table of Contents (TOC) O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=106 SRC="FIGDIR/small/737789v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@1ad85b2org.highwire.dtl.DTLVardef@1248bd4org.highwire.dtl.DTLVardef@58268dorg.highwire.dtl.DTLVardef@14b2da0_HPS_FORMAT_FIGEXP M_FIG C_FIG

12
Multiparameter optimization extends the lifetime of cell-free protein synthesis in a high-throughput format

Bozkurt, E. U.; Zanchet, B.; Nikel, P. I.; Volke, D. C.

2026-07-15 molecular biology 10.64898/2026.07.15.738603 medRxiv
Top 0.1%
2.8%
Show abstract

Cell-free protein synthesis (CFPS) is a powerful platform for synthetic biology, yet the factors governing reaction longevity remain poorly understood despite their importance for high-throughput applications. Here, the three principal determinants of CFPS performance--DNA template design, reaction composition, and lysate genotype--were systematically optimized to extend reaction lifetime in a 384-well plate format. Different energy regeneration systems were evaluated through real-time pH monitoring and metabolomic analyses to identify the metabolic constraints limiting prolonged protein synthesis. Lysates prepared from engineered Escherichia coli BL21(DE3) strains were further examined to assess the contributions of DNA, RNA, and amino acid stabilization. Systematic optimization of amino acid, nucleoside triphosphate, polyethylene glycol, and lysate concentrations identified DNA template stability and amino acid preservation as the primary factors sustaining CFPS activity. Combining these improvements yielded reactions that remained productive for >14 h and produced 567 {+/-} 64 g mL-1 active deGFP. These findings establish practical strategies for extending CFPS lifetime and improving high-throughput cell-free platforms.

13
Machine learning guided cell-free expression maps the biochemical landscape of carbonic anhydrase

Lazar, J. T.; Komp, E.; Martinez, I.; Zolkin, K.; Notin, P. M.; Saleh, S.; Landwehr, G.; Kim, K.; Tian, A.; Shapero, B.; Karim, A. S.; Marks, D.; Beckham, G. T.; Jewett, M. C.

2026-07-08 synthetic biology 10.64898/2026.07.07.736810 medRxiv
Top 0.1%
2.2%
Show abstract

Carbonic anhydrases are among the fastest known biocatalysts, reversibly facilitating the hydration of CO2 to HCO3- at rates up to 107 s-1, which warrants their investigation for industrial carbon capture technologies. However, engineering carbonic anhydrases to maintain stability under harsh industrial process conditions remains a key challenge, and sequence-to-function datasets compatible with machine learning to inform forward engineering are lacking. Here, we developed a high-throughput platform that couples cell-free gene expression with a gaseous CO2 colorimetric assay to map the fitness landscapes of carbonic anhydrases. From 96 diverse natural homologs, we identified a robust variant from the Aquificota phylum and conducted an exhaustive mutational scan and functional assessment of this enzyme at 70C and 90C, covering >99% of all single-amino acid substitutions (totaling 4,365 mutations assayed in 39,285 reactions). This biochemical landscape was used to benchmark 22 zero-shot protein fitness models and identify critical mutations that improved enzyme stability at 90C by more than three-fold. We then used both zero-shot protein language models and supervised learning to filter 419 model-generated variants from a ProteinMPNN library of 100,000 sequences, leading to a best-in-class enzyme that retained activity after incubation at 95C. This work demonstrates that integrating cell-free enzyme engineering with machine learning enables opportunities for high-throughput experimental measurements to benchmark and improve protein language models, accelerate design loops, and expand functional exploration within protein families where experimental information is limited.

14
Hierarchical Cytochrome P450 Oxidations Program Persiathiacin Assembly

Sumang, F. A.; Stevens, M. T.; Britton, W. J.; Errington, J.; Dashti, Y.

2026-07-09 microbiology 10.64898/2026.07.09.737402 medRxiv
Top 0.1%
2.1%
Show abstract

Thiopeptides are ribosomally synthesized and post-translationally modified peptides (RiPPs) that form complex bioactive scaffolds through extensive enzymatic tailoring. The polyglycosylated thiopeptides persiathiacins, exhibit potent activity against multidrug-resistant Mycobacterium tuberculosis (Mtb) and methicillin-resistant Staphylococcus aureus (MRSA). The persiathiacin biosynthetic gene cluster encodes six cytochrome P450 (CYP) enzymes, but the logic of their oxidative modifications was unknown. Here, we establish a protoplast-based genetic system for Actinokineospora and systematically assign functions to all P450s. We demonstrate that PerX hydroxylates the central thiazole, PerV installs the third indole-core crosslink required for macrocyclization, and PerT, not PerU, catalyses indole N-hydroxylation. Combined gene inactivation and metabolite profiling reveal a hierarchical enzymatic sequence leading to the mature scaffold prior to sugar installation. Notably, the intermediate accumulating in the {Omega}perX mutant exhibits enhanced anti-M. tuberculosis potency compared to persiathiacin A (IC50 = 0.07 vs 1.5 g mL1). These results define the enzymatic logic and temporal organization of persiathiacin biosynthesis, providing a conceptual framework for rational diversification of complex thiopeptide natural products.

15
EZSolver: Template-free prediction of polar enzymatic mechanisms via bidirectional flow matching and search

Kuo, L.-H.; Yang, J.; Arnold, F.

2026-07-09 bioinformatics 10.64898/2026.07.08.737313 medRxiv
Top 0.1%
1.9%
Show abstract

Predicting enzymatic reaction mechanisms is critical for understanding enzyme function and for designing and dis-covering new enzymes. Current computational predictors rely on deterministic, rule-based dictionaries, which per-form well on in-distribution tasks but fail to generalize to out-of-distribution (OOD) chemistry. To address this limita-tion, we present EZSolver, a template-free, generative framework for polar enzymatic mechanism prediction. Powered by a flow matching predictor (EZFlow) and navigated by an evaluator-guided bidirectional beam search, EZSolver learns the chemistry of electron redistribution instead of memorizing rigid templates. Evaluated across diverse en-zyme classes, EZSolver achieves a 60.0% accuracy and an 84.6% chemical plausibility rate for full mechanism predic-tion of unseen polar enzymatic reactions. While rule-based models collapse without predefined templates, EZSolver successfully extrapolates chemical knowledge to infer uncatalogued pathways, as demonstrated during rigorous OOD benchmarking. By illuminating enzymatic chemical mechanisms, EZSolver helps pave the way for automated predic-tion of enzyme function and discovery and design of novel biocatalysts for sustainable chemistry.

16
AI-enabled rhodopsin design for blue-light enhanced bacterial growth

Saeed, H.;Lewis, M.;Fujiwara, T.;Huang, J.;Konno, M.;Mori, K.;Yoshizawa, S.;Inoue, K.;Pan, T.;Wang, Y.;Yang, A.;Huang, W.

2026-06-30 Synthetic Biology 10.64898/2026.06.29.735265 medRxiv
Top 0.2%
1.2%
Show abstract

We developed an AI-guided design pipeline that generated and validated non-natural microbial rhodopsins with spectral properties not yet known in nature. The pipeline comprised a three-stage in silico design, a genetic algorithm (GA) for sequence generation, a stacked LASSO and XGBoost machine-learning (ML) regressor for spectral prediction and fitness ranking, and a Markov-based sequence plausibility filter to enforce proton pumping like characteristics. Four candidate rhodopsins (APR1, APR2, APR6, and APR7) targeting blue light absorption were designed and AlphaFold3 structural modelling predicted retinal binding pocket architecture consistent with outward proton-pumping function. Experimental characterisation confirmed that all four variants absorbed light at [~]410 nm and significantly promoted the growth of Cupriavidus necator under blue light illumination. This study demonstrates that AI-enabled design can engineer proteins with no natural precedent, generating light-harvesting rhodopsins with novel spectral properties while preserving biological function, marking a significant advance in programmable synthetic biology.

17
ThermoFusion: A Multimodal Deep Learning Framework for Generalizable Prediction of Enzyme Thermostability

Wei, Y.; Eberini, I.; Meyer, F.

2026-07-07 bioinformatics 10.64898/2026.07.04.736494 medRxiv
Top 0.2%
1.1%
Show abstract

Protein thermostability is a critical property for both industrial and biomedical enzyme applications, yet experimental evaluation of mutation-induced stability changes remains laborious and costly. Here, we present ThermoFusion, a hybrid deep learning framework that integrates 3D protein structure embeddings from ThermoMPNN with sequence-based embeddings from the pretrained protein language model ESM2 to predict the effects of single-point mutations on protein stability ({Delta}{Delta}G). ThermoFusion exhibits robust generalization, maintaining high predictive accuracy across out of distribution sequences with low identity to the training set -- a scenario where many other machine learning models, including ThermoMPNN and state-of-the-art tools, perform poorly due to reliance on memorization. Benchmarking on a curated enzyme dataset comprising of 105 enzymes and 3144 mutations shows that ThermoFusion reliably identifies stabilizing mutations while accurately predicting stability for enzymes beyond its training set. These results establish ThermoFusion as a powerful tool for rational enzyme design beyond its training set.

18
Genetic Code Expansion for Site-Specific Encoding of a Switchable, Intrinsic Fluorophore-Quencher Pair to Monitor Protein Dynamics

Giri, P.; Yarra, V.; Mathis, M.; Hurley, C.; Jones, C.; Eteme, O. N.; Hostetler, Z.; Cooley, R. B.; Kohli, R.; Mehl, R.; Petersson, E. J.

2026-06-29 biochemistry 10.64898/2026.06.26.734876 medRxiv
Top 0.2%
1.1%
Show abstract

Precisely modifying proteins at multiple sites in their native, folded structures offers unique opportunities to answer molecular and cellular-level biological questions. Here, we present a genetic code expansion strategy for site-specific integration of a fluorophore-quencher pair comprising two non-canonical amino acids--acridonylalanine (Acd) and methyltetrazinyl phenylalanine (Tet) -- into a protein expressed in E. coli. The Acd and Tet pair requires no post-translational labeling, and quenching can be switched off by biorthogonal or photochemical reactions of Tet for convenient internal control experiments. Mechanistic studies based on Stern-Volmer quenching, fluorescence lifetime measurements, and "proline ruler" peptides established the distance dependence of quenching. As proof-of-concept, we applied this strategy to study: 1) calmodulin, a calcium-sensing protein, 2) RecA, a DNA damage sensor in bacteria, and 3) LexA, a transcriptional repressor whose activation by RecA governs acquired antibiotic resistance in bacteria. Using these proteins, we demonstrate that dual Acd/Tet labeling provides molecular-level insights into protein dynamics, enables high-throughput drug screening, and advances tools for studying protein structure-function relationships.

19
Structural Determinants of Catalytic Directionality in an AMP-Forming Acetyl-CoA Synthetase from Syntrophus aciditrophicus

Yaghoubi, S.; Dinh, D. M.; Thomas, L. M.; Wofford, N. Q.; McInerney, M. J.; Follmer, A. H.; Karr, E. A.

2026-07-07 biochemistry 10.64898/2026.07.06.736832 medRxiv
Top 0.2%
1.0%
Show abstract

Acetyl-coenzyme A (CoA) is a central metabolic intermediate that links carbon and energy metabolism across all domains of life. The conversion of acetate and acetyl-CoA is carried out by three enzyme pathways: acetate kinase/phosphotransacetylase, ADP-forming acetyl-CoA synthetase, and AMP-forming acetyl-CoA synthetase (Acs). Acs enzymes serve critical physiological roles across diverse organisms generally by catalyzing a reversible two-step reaction forming acetyl-CoA and AMP from acetate and ATP. Isolated from the wastewater reclamation facility in Norman, Oklahoma, Syntrophus aciditrophicus strain SB (Sa) relies on an AMP-forming acetyl-CoA synthetase (SaAcs1) that favors synthesizing acetate and ATP from acetyl-CoA and AMP, in contrast to all previously characterized Acs enzymes. The origin of this preference and the structural determinants of both the thioester-forming step and catalytic directionality remain poorly understood. Here, we report a 2.2 [A] crystal structure of full-length SaAcs1 in the adenylation conformation with acetyl-AMP bound in the active site. Structural comparison to the extensively characterized Acs enzymes from Salmonella enterica (SeAcs) and Cryptococcus neoformans (CnAcs) revealed a displaced CoA-binding loop in SaAcs1. Enzymatic assays confirmed that SaAcs1 preferentially catalyzes the ATP-forming reaction. Site-directed mutagenesis demonstrated that reversion of two residues, G196 and T197, at the beginning of the CoA-binding loop to the consensus sequence repositions the loop and shifts catalytic preference toward the AMP-forming direction. Together, these results establish the CoA-binding loop and G196 and T197 as the primary structural determinants of directional preference in SaAcs1.

20
Site-Specific Introduction of Non-Canonical Amino Acids into natural and engineered Non-Ribosomal Peptides

Schreiber, M.; Dehghan, M.; Kibet, S.; Tvilum, M.; Kegler, C.; Hoffmann, K.; Gruen, P.; Balluff, S.; Siems, K.; Bode, H. B.

2026-07-13 biochemistry 10.64898/2026.07.12.738027 medRxiv
Top 0.2%
1.0%
Show abstract

The incorporation of non-canonical amino acids (ncAAs) into proteins, developed in the past 20 years, has opened new avenues with respect to protein structure, protein modification, protein-protein interaction or enzyme catalysis beyond what is possible with the 20 proteinogenic AAs. Although >300 unusual building blocks including several ncAAs have been described in nonribosomal peptides (NRPs) naturally, we aimed to further expand the scope of the underlying nonribosomal peptide synthetases (NRPS) to incorporate ncAAs beyond the naturally available ones. We have therefore systematically screened for ncAA accepting NRPS systems, applied NRPS engineering to transfer the respective ncAA-accepting parts into other NRPSs and thereby created novel peptides that were further derivatized in post-enzymatic chemical synthesis reactions directly in bacterial culture extracts. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=177 SRC="FIGDIR/small/738027v1_ufig1.gif" ALT="Figure 1"> View larger version (38K): org.highwire.dtl.DTLVardef@90552forg.highwire.dtl.DTLVardef@1c8a5e0org.highwire.dtl.DTLVardef@2549dorg.highwire.dtl.DTLVardef@1012911_HPS_FORMAT_FIGEXP M_FIG C_FIG